Medical Decision Making
○ SAGE Publications
Preprints posted in the last 30 days, ranked by how well they match Medical Decision Making's content profile, based on 12 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Scherer, L. D.; Matlock, D. D.; Cronin, J.; Gritz, M.
Show abstract
Multi-Cancer Detection (MCD) tests can detect more than 50 different types of cancer using a blood test. Recently passed law in the U.S. guarantees that Medicare will pay for these tests when they are FDA approved and show evidence for clinical benefit. This manuscript provides estimates of the cost of MCD tests to Medicare under different assumptions of cost per test, eligibility, and screening uptake in the eligible population. This manuscript additionally estimates the cost of follow-up testing resulting from false positive results, which are considered avoidable costs caused by the screening test.
Maleki, C.; Bertrand, Y.; Gailly, F.
Show abstract
Clinical recommendations are often expressed in narrative form, which limits their direct execution, auditability, and patient-specific interpretation. This paper presents a hybrid decision-support framework that combines Decision Model and Notation (DMN), survey-weighted rule-ensemble learning, and counterfactual sensitivity analysis. The framework is evaluated using an NHANES-derived fasting cohort for classification of documented diabetes status. The full fasting analysis cohort contained 2,582 participants, and a non-diagnostic laboratory subgroup, Gate0, contained 2,111 participants. On untouched test data, the rule-ensemble model achieved ROC-AUC and PR-AUC values of 0.959 and 0.873 in the full fasting cohort and 0.861 and 0.499 in Gate0. Four clinically interpretable candidate rules were selected using validation data only. A nonnegative survey-weighted logistic model removed one redundant rule and converted the remaining three binary activations into an auditable DMN score and model-estimated probability. The final DMN achieved ROC-AUC 0.769, PR-AUC 0.153, and Brier score 0.029 in the untouched Gate0 test set. In small rule-defined test subgroups, hypothetical five-unit BMI reductions lowered mean model-estimated probability by 2.40 to 5.89 percentage points when one or more BMI thresholds were crossed. These findings characterize policy sensitivity rather than causal effects and require external validation.
Ohno, K.; Hirai, M.; Hashimoto, S.
Show abstract
Background: Descriptive mapping of intensive care unit (ICU) and high care unit (HCU) capacity across Japan's secondary medical areas (SMAs) characterizes where beds exist, but medical planning also requires answers to prospective questions: how likely is a capacity shortfall under demand surge, which assumptions drive that risk, and how much protection do inter-zone transfer arrangements provide. No openly available tool addresses these questions at the SMA level, the geographic unit at which Japanese medical plans are written. Methods: We developed MeshScope-Scenario, a probabilistic capacity-demand framework operating on the MeshScope-Region platform. For a selected SMA, observed inputs (notified ICU/HCU beds from the Hospital Bed Function Reports; resident population) are combined with four explicitly flagged assumption parameters - effective staffed-bed rate, concurrent severe-care demand per 100,000 population, surge multiplier, and net cross-boundary inflow - each with a user-specified distribution. A seeded Monte Carlo engine (deterministic reproduction under a fixed seed) estimates the distribution of bed shortfall; interventions are compared under common random numbers. Parameter dependence is introduced by a Gaussian copula with automatic positive-semidefinite correction; global sensitivity is quantified by Sobol first-order and total-order indices (Saltelli sampling, Jansen estimators) alongside a deterministic one-at-a-time tornado analysis. A two-zone extension transfers unmet demand to the nearest ICU-holding SMA using road-network travel times measured in MeshScope-Region, with a transfer time limit and an acceptance cap; because both zones share the same systemic draws, correlated exhaustion of donor capacity under surge ("shared-fate" risk) is represented structurally. A seasonal layer applies twelve monthly surge multipliers and reports the distribution of annual maximum shortfall and month-specific shortfall probabilities. The engine is a dependency-free pure-function module verified by 40 statistical tests. Results: The framework reproduces identical output under identical seed and input; a flat seasonal profile reproduces the non-seasonal model exactly; copula factorization error is below 1e-9; and 10,000 iterations across three intervention variants complete in approximately 50 ms in a standard browser, permitting fully interactive use. Applied to three archetypal SMAs from the observed FY2024 supply map (seed 42, 20,000 iterations, demand prior 5 per 100,000), an ICU-zero zone with a transfer partner 48 minutes away has shortfall probability 66.0% (P50 4.8, P90 20.7 beds); a median metropolitan zone, 32.7% (P90 5.9), with Sobol indices ranking demand density and surge dominant; a high-supply zone shows zero shortfall up to approximately 2.5x surge. For the ICU-zero zone, independent-donor reasoning credits the transfer arrangement with a 2.27-bed reduction in expected shortfall, of which shared-fate correlation removes 93%; the probability of severe shortfall under the arrangement equals that with no arrangement at all, while a committed pool of five donor beds (6% of donor effective supply) restores a 13-point reduction and more than halves the arrangement's correlation exposure. Conclusions: MeshScope-Scenario extends SMA-level capacity mapping from description to prospective risk assessment. All demand-side inputs are declared assumptions with adjustable distributions rather than estimates presented as fact; the framework's value is to make the consequences of those assumptions, and their interaction with observed supply, explicit, reproducible, and inspectable for planning deliberation.
Brodsky, S.; Matlin, O.
Show abstract
Improving primary care is a long-standing strategy to constrain health care spending. Yet, evaluations of primary care models focused on payment reform have shown minimal effects on total cost of care. We report the results from a large-scale, real-world evaluation of an advanced primary care model that restructures access through same-day and next-day appointments, on-demand video visits, asynchronous clinician messaging, and extended hours. Using a stacked-cohort difference-in-differences design with entropy balancing and inverse probability of censoring weighting, we analyzed multi-payer claims covering April 2022 through March 2025. Advanced primary care use was associated with an 8.6% reduction in total cost of care (-$729 per patient per year; P = 0.004), driven by lower specialist cost (-$939/year; P < 0.001) and, to a lesser degree, by reductions in inpatient (-$134/year; P < 0.001), urgent care (-$70/year; P < 0.001), and emergency department cost (-$16/year; P = 0.02), partially offset by higher primary care cost (+$350/year; P < 0.001). The specialist reduction was concentrated in knowledge-based consultative encounters (-$663/year; P < 0.001), while procedural specialist cost was largely unchanged (-$276/year; P = 0.09). Cost differences emerged in the first post-index month. These findings suggest that advanced primary care may reduce total health care spending, with observed savings driven primarily by lower spending on consultative specialty care.
Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.
Show abstract
Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.
Zanwar, P. P.; Wang, M.; Logan, N.; Chang, S.-H.
Show abstract
Introduction: Research has documented that obesity and morbidity are associated. Black persons in the United States (U.S.) incur higher financial costs of obesity-related multimorbidity (ORM). However, lifetime healthcare costs (LHCs) remain underexamined for these populations. Objective: We quantified racial differences in 1) LHCs and 2) lifetime healthcare cost differential (LCD) associated with ORM for ages > 40 years. Methods: We used the 2008- 2012 Medical Expenditure Panel Survey Household Component to examine unique obesity-related diseases (ORDs): high blood sugar, hypertension, coronary heart disease, and stroke. We used a prior published Markov model to simulate a person's life history of ORDs and compute LHCs among ages > 40 years. We computed LCD-associated ORM as the difference in LHC for those with ORM and LHC for members without ORDs. We quantified differences in race as the difference between LHC or LCD among White and Black men and women. Results: Our analytic sample included 53,035 Black and White persons representing 97,229,611 (S.E., 2,104,365), 12.4% as Black and 87.6% as White persons. ORM was more prevalent in the Black (21.2%) than the White group (13.4%). LHCs by race (Black/White) for women/men with ORM and LCDs associated with ORM (2012$) were $3 1,035/43,595 and $11,350/26,948 for age 40-49, $2 1,567/25,6 115 and $3,846/9,808 for 50-59, $9,863/18,515 and -$2,566/7,426 for 60-69, -$8,220/16,285 and -$11,524/3,865 for 70-79. Conclusions: Racial Differences in LHCs and LCDs related to ORM persist and vary across subpopulations. Future interventions designed to prevent/manage ORM are crucial for prioritizing populations with high LHCs and advancing health equity.
Aldis, R.; Wang, S.; Sage, M.; Metzmaker, M.; Galvin, H.
Show abstract
Ambient artificial intelligence scribes are being increasingly used in healthcare to improve efficiency and reduce provider clinical documentation burden, yet their performance across linguistically diverse patient populations is not well characterized. We conducted a retrospective analysis of 54,160 outpatient encounters within a U.S. safety net health system to evaluate the performance of an artificial intelligence documentation tool in English and non-English clinical encounters, and in encounters where an interpreter or bilingual provider was present. Documentation performance was measured by the percentage of words in the final note that were generated by the ambient AI documentation tool and not edited by the provider. Associations between language factors and documentation performance were measured using Generalized Estimating Equations with exchangeable correlation structures to account for clustering of multiple encounters within unique patients. Univariable models were fitted to estimate the odds of adequate performance by language and interpreter modality, and a multivariable interaction model was used to evaluate within-language differences between bilingual providers and interpreter-mediated encounters. Non-English encounters were 21% to 25% less likely than English encounters to achieve the same performance threshold. There was no significant difference in generative documentation performance between interpreter-mediated and bilingual provider encounters. These findings underscore the importance of equity-focused evaluation and multilingual model refinement to ensure that artificial intelligence documentation benefits are distributed fairly across diverse patient populations.
Esteban, S.; Quintana, G.; Sanchez, M.; Szmulewicz, A.
Show abstract
Background: Digital reminders reduce outpatient no-shows, but the optimal timing and frequency of messages remain unclear, particularly in Latin American public health systems. We emulated a target trial to evaluate the comparative effectiveness of four WhatsApp reminder strategies on appointment absenteeism and patient-initiated cancellations. Methods: We analyzed administrative and electronic health-record data from the public health system of the Autonomous City of Buenos Aires, Argentina (June 2023-May 2024). Eligible individuals had scheduled an in-person outpatient appointment in one of 15 prioritized specialties at least 75 hours in advance and had a mobile phone on record. We compared four strategies: (1) dual reminders at ~72 and ~24 hours before the appointment; (2) a single reminder at ~72 hours; (3) a single reminder at ~24 hours; and (4) no reminders. The primary outcome was the proportion of no-shows by the end of follow-up. Secondary outcomes were the cumulative incidence of patient-initiated cancellations overall, within 12 hours of the appointment, and followed by rebooking. We emulated the target trial using a cloning-censoring-weighting approach to estimate per-protocol controlled direct effects, with inverse-probability weights to address time-varying confounding and selection bias. Cumulative incidence of secondary outcomes was estimated using weighted Kaplan-Meier curves. Three pre-specified sensitivity analyses and standardized mean differences assessed robustness and covariate balance. Results: A total of 475,214 first eligible person-appointments were included; baseline no-show risk in the control arm was 34.6%. All three active strategies reduced no-shows compared with no reminders. The single 24-hour reminder produced the largest reduction (Risk Ratio [RR] 0.76, 95% CI 0.72, 0.81; Risk Difference [RD] -8.21 percentage points [pp], 95% CI -9.68, -6.54), followed by the dual-reminder strategy (RR 0.80, 95% CI 0.79,0.81; RD -7.05 pp, 95% CI -7.41, -6.71) and the single 72-hour reminder (RR 0.91, 95% CI 0.84,0.99; RD -3.16 pp, 95% CI -5.69, -0.49). All active strategies increased patient-initiated cancellations relative to control, with the dual-reminder strategy producing the largest increase. Sensitivity analyses preserved the qualitative ranking of strategies across all specifications. Conclusions: In this large target trial emulation, a single just-in-time WhatsApp reminder sent ~24 hours before the appointment was as effective as a dual-reminder schedule in preventing no-shows and superior to a distal 72-hour reminder alone. Adding a second, distal reminder provided no measurable benefit for attendance but substantially increased patient-initiated cancellations, which may be operationally valuable when active slot reallocation is a goal. These findings support timing, rather than frequency, as the primary lever of digital-reminder effectiveness, and favor the deployment of a single proximal reminder as the default strategy in resource-constrained outpatient settings.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
Malhotra, A.; Patel, N.; Kaftan, D.; Mudimu, E.; Bershteyn, A.; Sharma, M.
Show abstract
Introduction: Monthly oral HIV pre-exposure prophylaxis (PrEP) such as MK-8527 offer promise as low-cost, self-administered options that are easily delivered through community-based platforms. Economic evaluations of MK-8527 either alone or alongside other long-acting (LA) PrEP products like lenacapavir are needed for informing HIV prevention strategies. Methods: We adapted an agent-based network model, EMOD-HIV, to simulate LA-PrEP scale-up in South Africa and western Kenya from 2026-2035; scenarios evaluated MK-8527 alone, lenacapavir alone and combined strategies, with varying uptake among female sex workers, their male clients, and individuals with >1 partner. We assumed 95% effectiveness of MK-8527 for 2 months (assuming individuals took 2 of 3 pills dispensed) and 95% lenacapavir effectiveness for 6 months. Scenarios were compared to a baseline of daily oral PrEP only. Results: Assuming the same uptake rates, MK-8527 alone achieved lower health impacts than lenacapavir alone in both settings (6-14% vs. 11-18% of HIV infections averted) but had substantially lower costs; provision costs of MK-8527 were 58-59% lower than lenacapavir assuming $1.00/pill and 40-43% lower at $2.50/pill. In western Kenya, ICERs for MK-8527 alone were $467/DALY averted and $799/DALY averted assuming pill prices of $1.00 and $2.50 respectively, compared to $1,306/DALY averted for lenacapavir. In South Africa, all LA-PrEP strategies were cost-saving over the 35-year horizon, although near-term budget impacts were substantial ($169-348 million over five years). Service delivery accounted for the majority of MK-8527 costs (73% at $1.00 per pill). Combined strategies of lenacapavir and MK-8527 increased health benefits (13-23% infections averted) but had higher provision costs than either strategy alone. Conclusion: MK-8527 can reduce HIV incidence at lower costs than lenacapavir. However, high service delivery costs limit its cost-effectiveness to scenarios in which pill prices are low (US$1.00 per pill) and provision is targeted to populations at substantial HIV risk.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Moe-Byrne, T.; Knapp, P.; Golder, S.
Show abstract
Background People with lower levels of literacy or health literacy may struggle to understand conventional health information. Video animations show promise as information tools, yet it is unclear whether video animations help reduce these inequalities in understanding. This study examined whether the effectiveness of video animations in health settings differs according to level of literacy or health literacy. Methods We drew on trials from a recent systematic review of video animations about healthcare or public health topics for patients or the public. We extracted available data on literacy, health literacy, or proxy indicators. One reviewer extracted data and a second checked all entries. Where possible, we conducted subgroup analyses of low and high literacy levels or interaction meta-analyses comparing low versus high literacy groups; otherwise, results were summarised narratively. Results From 88 eligible trials, we extracted health literacy data for 12. Across nine trials reporting knowledge, animations mostly improved knowledge compared with controls in both lower and higher health literacy groups. Effects on attitudes and behaviours were mixed and often small, with few studies reporting results by health literacy level. Across the subgroup analyses available, there was no consistent evidence of a pooled interaction effect of animations according to low and high literacy groups, but both statistical heterogeneity and small subgroup sizes limited precision of estimates. Across 88 trials, 54 (61%) reported education level, 22 (25%) did not, and 12 (14%) involved children or adolescents likely to have similar education levels. Conclusions Overall, the available data suggest that video animations can improve knowledge outcomes in both lower and higher health literacy groups, but their impact on attitudes and behaviour is less clear. Because literacy was rarely reported or analysed in the trials, it remains uncertain whether animations help to reduce literacy-related inequalities in access to, and use of health information.
Gorenshtein, A.; Omar, M.; Barash, Y.; Kruskal, J. B.; Ahmed, M.; Brook, O. R.; Klang, E.
Show abstract
Clinical AI agents may be assigned to individual patients, but hospital resources are shared across many patients. We tested what agents do when helping their assigned patient would violate the hospital's rule for a scarce resource. We analyzed 22,916 simulated cases comprising 274,992 logged agent actions across 20 AI models. In each scenario, the agent could claim a scarce resource for its patient even though the hospital rule gave another patient priority. We varied only the agent's assigned role, from responsibility for the whole ward to strong advocacy for one patient. Violations of the hospital rule rose from 32.5% under whole-ward responsibility to 69.4% under strong patient advocacy, a 36.9-point increase (95% CI, 25.7-48.0). Agents correctly identified which patient should receive the resource in 95.7% of tests, yet still took it for their own patient in 65.9% of those episodes. Asking the agent to apply its own allocation judgment immediately before acting reduced violations to 0-2% in a three-model follow-up experiment. Assigned roles can shape how clinical AI agents use shared hospital resources, even when they identify the correct priority patient. Patient-focused agents should not independently control shared resources without an allocation check.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Yang, C.-H.; Salvatore, M.; Lu, H.; Zhu, Z.; Tennant, P.; Shi, X.; Ohno-Machado, L.; Khera, R.; Gross, C.; Li, F.; Mukherjee, B.
Show abstract
Electronic health record (EHR)-linked cohorts support association, prediction, and causal studies using longitudinally measured markers of health. However, a lab biomarker measurement is recorded only when a patient first has a medical encounter (visit process) and, a clinician orders the corresponding test and the patient follows through (observation process). These two stages may induce informative presence (IP) and informative observation (IO), respectively. Yet their drivers remain largely uncharacterized, despite evidence that understanding this recording mechanism is essential for selecting appropriate strategies for downstream analysis that treat these markers as longitudinally measured outcomes. We characterize this two-stage recording hierarchy using a stochastic recurrent-event model for the outpatient visit process and a visit-process-weighted generalized estimating equation model for biomarker recording conditional on an outpatient visit. We characterize descriptors of both processes in three EHR-linked cohorts in the US (All of Us [AoU], n=599,423; Yale New Haven Health System [YNHHS], n=319,666; Michigan Genomics Initiative [MGI], n=82,372), reporting descriptive statistics for longitudinal visits and for a panel of 68 lab biomarkers commonly measured in EHRs. We conduct detailed model-based analyses of ten biomarkers spanning multiple domains: routine monitoring, general laboratory assessment, and symptom-triggered testing. These include glucose, hemoglobin A1c [HbA1c], creatinine, hemoglobin [Hgb], white blood cell count [WBC], low-density lipoprotein [LDL] and high-density lipoprotein [HDL] cholesterol, triglycerides, C-reactive protein [CRP], and thyroid-stimulating hormone [TSH]. Across the three cohorts, the median number of outpatient visits ranged from 1.7 to 6.1 per year over a median follow-up of 4.4 to 7.2 years. Among patients with at least one recorded measurement, the median within-person proportion of visits containing a given biomarker ranged from 0.4% to 19.5%, demonstrating that more frequent visits did not necessarily translate into greater per-visit biomarker capture. In the visit-process models, chronic disease burden, and a recent history of outpatient visits were consistently associated with higher visit rates across all three cohorts whereas associations with race, ethnicity, and neighborhood-level income varied across cohorts. In per-visit observation models, the association of covariates depended on the biomarker under consideration; for example, prior cancer diagnosis was associated with more frequent measurement of blood counts but with less frequent measurement of lipids. These findings provide a deeper understanding of how to model who seeks care and what is measured as two distinct recording processes in EHR. Our empirical findings show that the descriptors of these processes vary across cohorts and biomarkers, providing guidance on how to construct these models for downstream longitudinal analyses with irregular EHR visits.
Smith, S. J.; Lemoine, D.
Show abstract
Objective: To assess the efficacy of an executive peer coaching program, Charting Champions Program (CCP), in helping physicians manage their administrative workload, thereby improving time management, workflow and well-being. Findings: In this longitudinal survey study, physicians self-reported significant improvements in completing charting and administrative paperwork during their clinical day. Physicians reported significant improvements in mental, cognitive and emotional states after the program. Meaning: The Charting Champions Program is an effective intervention that supports physicians in problem-solving the administrative burden of their clinical day, improving workflow efficiency, completing administrative requirements during clinical hours, and enhancing work-life balance and personal satisfaction. Background: Physicians are subject to high levels of mental, physical, and emotional stress, partly due to increasing administrative burdens. Online coaching is a proven intervention to help physicians improve workflow efficiency, reduce administrative burden and improve job satisfaction. Design: This voluntary longitudinal survey took place between 2020 and 2023. Physicians were asked to complete a survey at program entry and again 30-90 days after program completion. The survey consisted of 14 Likert scale questions, and a final sample of 280 physicians completed both surveys. Intervention: CCP contains modules that teach workflow improvements for clinical days, including timely charting, administrative task workflow, managing patient consultations and reducing interruptions. Interventions include self-paced modules, live coaching, recordings and an online peer community. Results: Post-CCP physicians reported a significant decrease in hours spent charting (P<0.0001) and completing clinical paperwork outside of clinical hours (P<0.006). Physicians also reported a decrease in work-related dread (P<0.001), feelings of burnout (P<0.001), and thoughts of quitting due to administrative burdens (P<0.001). Physicians felt more focused at work (P<0.001), felt more in control of the clinical day (P<0.001), and rated their mental energy at work higher (P<0.001). The program did not affect the number of patients seen in a full clinical day (P > 0.918). Conclusion and Relevance: The CCP reduces the time physicians spend on tasks outside of clinical hours, increasing free time without decreasing the number of patients seen per day.
Yano, Y.; Nagasu, H.; Hiroshi, K.; Ohashi, M.; Isaka, Y.; Okada, H.; Nangaku, M.; Kashihara, N.
Show abstract
Background: Traditional real-world studies comparing SGLT2 and DPP4 inhibitors on renal outcomes rely on propensity score matching, which causes high-dimensional data loss. We used causal machine learning (Causal ML) to unmask heterogeneous treatment effects in diabetic kidney disease (DKD). Methods: Using data from 4,588 patients within the Japanese J-CKD-DB-Ex registry, we implemented a doubly robust (DR) learning framework (Linear DR-learner with XGBoost) to compare SGLT2 and DPP4 inhibitors. Outcomes included the chronic eGFR slope and a composite renal endpoint ([≥] 50% eGFR decline or end-stage kidney disease). Heterogeneity was explored via causal SHAP and decision trees. Results: At the population level, SGLT2 inhibitors modestly slowed chronic eGFR decline (average treatment effect [ATE] = 0.14 [95% CI: -0.86, 1.15] mL/min/1.73m^2/year) and reduced composite endpoint risk by 9% (ATE: -0.09 [-0.11, -0.08]) versus DPP4 inhibitors. However, individual-level counterfactual analysis suggested that for the chronic eGFR slope, non-glinide users with stable pre-treatment trajectories who were also taking ACE inhibitors had a greater benefit from SGLT2 inhibitors (ATE: 2.95 [-0.68, 6.58]). Conversely, glinide users with steep pre-treatment decline had a greater benefit from DPP4 inhibitors (ATE: -8.98 [-16.11, -1.85]). For composite renal events, SGLT2 inhibitors had a 28% absolute risk reduction within the algorithmically identified high-risk subgroup (eGFR [≤] 28.1 mL/min/1.73 m^2 and positive proteinuria; ATE: -0.28 [-0.33, -0.23]). Even non-proteinuric decliners demonstrated a 8% risk reduction with SGLT2 inhibitors (ATE: -0.08 [-0.10, -0.06]). Conclusion: Causal ML advances precision medicine in DKD, shifting from uniform prescribing to individualized, data-driven therapy targeting distinct intrarenal pathways.
Biglarbeigi, P.; Dale, C.; Lambarth, A.; Mason, A.; Takher, R.; Ballabio, G.; Minshull, J.; Mamas, M. A.; Tomlinson, C.; Rowark, S.; Rayman, G.; Pearson, E. R.; Khunti, K.; Sattar, N.; Sofat, R.
Show abstract
Objectives: To examine the conformance to type 2 diabetes NICE guidelines across cardiovascular risk strata; and to quantify geographical variation in treatment pathways following the COVID-19 pandemic, encompassing guideline changes. Design: We carried out a retrospective observational study using linked electronic health records across England. Process mining, a data driven method that can reconstruct clinical treatment pathways, was applied to map 12-month treatment trajectories after treatment initiation. Conformance with NICE NG28 (2022) was quantified using a structural similarity index. Further, behavioural and entropy-based similarity (capturing treatment variability and complexity) measures were used to assess sequencing and heterogeneity of treatment. Setting: Primary and secondary care in England datasets within the National Health Service England Secure Data Environment (NHSE SDE), analysed first at national level and then across 42 Integrated Care Boards (ICBs) which are the devolved health care geographical delivery regions in England. Participants: 822,650 individuals with newly diagnosed T2DM between 1-February-2022 and 1-November-2025, stratified into low cardiovascular risk (LR-C; QRISK3<10), high risk (HR-C; QRISK3>=10 or receiving statins/blood pressure lowering treatment), and established cardiovascular disease (eCVD-C). Participants were followed for 12 months after first dispensed glucose lowering therapy. Main outcome measure: First line therapy, treatment intensification and switching within 12 months; change in glycated haemoglobin (HbA1c); quantified conformance to NICE recommended pathways; and regional variation in broader similarity measures. Results: Metformin monotherapy was the dominant initiation strategy in LR-C and HR-C cohorts (92.4% and 90.2%, respectively), whereas eCVD-C showed lower uptake of metformin (68.9%) and higher uptake of SGLT2 inhibitors (26.3%). Intensification from metformin to combination therapy was infrequent across all cohorts (<1%), although HR-C demonstrated the highest treatment transitions and switching behaviour. Dispensed SGLT2 inhibitor use was nearly threefold higher in eCVD-C (26.7%) than in LR-C (9.0%) or HR-C (10.7%). Overall, conformance to NICE-recommended pathways remained modest nationally, particularly in LR-C and HR-C. Across 42 ICBs, substantial regional heterogeneity in treatment pathways and guideline conformance was observed, with conformance ranging from 0.29 to 1.00 in LR-C pathways, 0.40 to 1.00 in HR-C pathways, and 0.54 to 0.92 in eCVD pathways. Conclusion: National T2DM treatment pathways for post-pandemic showed higher alignment to NICE guidelines in eCVD-C compared to the other risk groups, with ongoing gaps and large regional variations in other risk groups. Process mining offers a scalable approach to monitor implementation of guideline recommended care that could support learning health systems. Using T2DM during and post COVID-19 pandemic as a case study, this work demonstrates how these methods can assess the use of existing and innovative therapies, identify gaps and guide future adoption to ensure recommended treatments reach the right patient groups.
Chaudhry, R.; Chen, Z. S.
Show abstract
Background: Mortality prediction models often combine early electronic health record data, but the relative prognostic value of baseline vulnerability, physiological severity, treatment exposure, and procedure burden remains unclear. Objective: To compare routinely available first-24-hour clinical domains for visit-level 30-day mortality prediction and assess whether domain-level patterns replicated in MIMIC-IV. Methods: We used CHoRUS, an OMOP-formatted acute-care dataset, with independent domain-level replication in MIMIC-IV. CHoRUS included 22,098 visits among 5,892 unique patients, with 1,004 30-day mortality events and 4.5% mortality prevalence. MIMIC-IV included 23,000 acute-care visits among 10,006 unique patients, with 819 events and 3.6% prevalence. Across both datasets, 45,098 visits and 15,898 unique patients were analyzed. Predictors were restricted to the first 24 hours after visit start. Performance was evaluated using AUPRC, AUROC, Brier score, calibration, sensitivity at 90% specificity, highest-risk 10% analyses, decision-curve analysis, and SHAP summaries. Because 30-day mortality was infrequent, the classification task was class-imbalanced. Accordingly, AUPRC was interpreted relative to the prevalence-based no-skill baseline, rather than as an absolute measure alone. Results and Conclusion: Physiological severity produced the largest improvement beyond baseline in CHoRUS, with median AUPRC 0.38 and median AUROC 0.86, and showed the same primary domain-level pattern in MIMIC-IV. Treatment exposure and procedure burden provided smaller gains. In CHoRUS, the best pairwise model combined baseline, physiological severity, and procedure burden features, with median AUPRC 0.41; the all-domain model was slightly lower, with median AUPRC 0.40 and median AUROC 0.86. In MIMIC-IV, the all-domain model had the highest median AUPRC, 0.25, only modestly above the best pairwise model. First-24-hour physiological severity features therefore provided the most consistent prognostic information across datasets, supporting parsimonious, clinically interpretable acute-care risk models centered on high-quality early physiological data.
Vijay, A.; Govind, N.; Moorthy, A.; Dunn, P.; Lababidi, Z.; Jones, S.; Stahlberg, M.; Ibrahim, S.; Koochek, K.; Shah, K. S.; Schulhauser, R.; Lerma, E. V.; Nair, L.; Livi, J.; Kalra, D. K.; Wadwekar, D.; Gulllett, W.; Vijayaraghavan, K.
Show abstract
Abstract Background: Cardiovascular-kidney-metabolic (CKM) syndrome is an increasingly prevalent multisystem condition associated with morbidity, fragmented care, recurrent hospitalization, and rising healthcare costs. While cardiovascular risk models estimate future disease risk, fewer frameworks support multidisciplinary CKM care, clinician decision-making, and population health management. Synthetic data environments can assess implementation readiness while preserving privacy. Methods: We validated the explainable, clinician-supervised CHAPERONE-CKM framework using a reproducible synthetic cohort of 10,090 simulated patients with 128 demographic, laboratory, imaging, treatment, and healthcare utilization variables across the CKM continuum. Synthetic data generation was separated from framework evaluation through probabilistic modeling and independent validation to reduce deterministic relationships. The framework generated CKM stage assignments, implementation priorities, clinician-readable rationales, multidisciplinary referral pathways, and guideline-directed therapy prompts. Evaluation focused on implementation readiness, consistency, calibration, subgroup stability, fairness, workflow simulation, and explainability. Results: The synthetic population represented CKM-related conditions including diabetes (52%), hypertension (65%), chronic kidney disease (20%), heart failure (32%), and prior CKM hospitalization (27%). The framework showed stable internal behavior across demographic and clinical subgroups, favorable calibration, and biologically plausible prioritization of advanced CKM disease. Workflow simulations suggested earlier identification of patients suitable for multidisciplinary review, therapy optimization, and coordinated care compared with reactive workflows. Traditional performance metrics supported framework behavior but were treated as secondary evidence rather than proof of clinical effectiveness. Conclusions: In a synthetic validation environment, the CHAPERONE-CKM framework demonstrated implementation readiness, transparent decision pathways, and compatibility with multidisciplinary CKM population health management. These findings are an early translational milestone, not clinical validation, and support external validation, prospective implementation studies, and Learning Health System integration to assess effects on care delivery, equity, and value-based outcomes.